Skip to content

feat(storage): route content binaries, metadata and cleanup through S3 asset storage - #37772

Open
swicken wants to merge 7 commits into
s3-stack/1b-binary-apifrom
s3-stack/2-content
Open

swicken wants to merge 7 commits into
s3-stack/1b-binary-apifrom
s3-stack/2-content

Conversation

@swicken

@swicken swicken commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

S3 asset storage, part 3 of 7. Stacked PRs, review bottom up. Each one builds on the one below.
1a storage layer #37770 · 1b binary asset API #37771 · 2 content #37772 · 3 recovery #37773 · 4 publishing #37774 · 5 temporary uploads and WebDAV #37775 · 6 rendering #37776
Everything is behind FEATURE_FLAG_S3_ASSET_STORAGE, off by default. With the flag off, behavior matches main.

Refs #37868

Proposed Changes

This is where content starts using S3, and the PR with the most changes to heavily used classes (ESContentletAPIImpl, Contentlet, FileMetadataAPIImpl, the JSON and XML serializers). Every change is gated.

  • Immutable revisions at check-in. New binaries are stored at .../{field}/.revisions/{revision-id}/{filename} and the Binary field JSON gains storageKey and metadataStorageKey. The reference changes inside the content transaction, so a rollback keeps the previous revision, and the revision a rolled-back check-in uploaded is deleted. Legacy JSON and keys stay readable.
  • Metadata. Generation restores evicted originals under a cache lease, byte-derived Tika extraction is shared across identical bytes (SharedExtractedMetadata), metadata keys follow the binary revision, and custom metadata is copied to replacements.
  • Durable deletion. Whole-inode deletion, per-language deletion and old-version cleanup record one binaryAssetCleanup job per inode in the deleting transaction. Each job records the exact paths stored at that moment and deletes only those, so a revision uploaded later under a reused inode (as a push-publishing receiver does) is never touched. Binary field deletion records binaryFieldCleanup jobs that archive, verify and then delete. Both use the existing job queue and its retry policy, and both reject direct submissions through the public job endpoint.
  • Recovery archives. With BACKUP_DELETED_CONTENTLETS_TO_DISK also on, deletion first writes a verified recovery ZIP to the deleted-content-backups group; a failed backup aborts the deletion.
  • Job processor registration filter. AssetStorageFeature.allowsJobProcessor keeps the S3 job processors out of JobQueueManagerAPIImpl while the flag is off.
  • Binary HTTP responses (BinaryExporterServlet), FileAsset.getInputStream and Contentlet.getBinaryStream hold a cache lease only until the file is open, so a slow client cannot defer eviction. That is safe on a local disk because an open file survives eviction; eviction on an NFS asset directory has not been validated.

Behavior with the flag off

Unchanged from main: no revision keys are written, no S3 cleanup processors are registered, and the deleteAllVersionsandBackup interceptor stays the existing no-op. The flag-off integration run below exists to prove this for the classes every check-in goes through.

Review fixes

The commit fix(storage): make whole-inode cleanup exact and reclaim rolled-back and per-language binaries addresses a full review of this PR. All of it is flag-on only:

  • Whole-inode cleanup deletes only the paths recorded at enqueue, instead of everything under the inode's prefix.
  • binaryAssetCleanup validates submissions the same way binaryFieldCleanup does.
  • A rolled-back check-in deletes the revision it uploaded (and that revision's metadata), logging rather than throwing if the delete fails. Metadata stored under a different, content-addressed key and savepoint rollbacks are not covered; the doc says so.
  • Deleting one language's versions now records cleanup jobs and the recovery archive, which main skips (Creating new page adds back reference to old contents #9146). To make those jobs commit with the deleted rows, ContentletAPI.delete(Contentlet, User, boolean) runs in a transaction with the flag on, because the Site Browser and WebDAV call it without one. This one goes beyond the review findings, so it is worth a look.
  • One cleanup job per inode, where duplicated version lists used to record several.
  • BinaryExporterServlet releases its lease once the file is open, as described above.
  • The unreachable flag-on branch in DropOldContentletRunner.deleteFromAssetsDir is removed.

Second review fix

A second review found that Contentlet.getBinary looked up a stored revision under the contentlet's current inode. A contentlet that carries the previous version's binary under a new inode, which is what check-in has after it assigns the new inode, threw File System error. once that revision had been evicted from the local cache, because the revision key names the old inode. The last commit (fix(storage): restore a carried-forward binary revision from the owner its key names) restores the revision under the owner its key names. BinaryAssetReference.localFile still validates the key. Flag-on only.

Check-in does not hit this today: saving the content JSON hydrates metadata, which restores the evicted file before check-in reads the binary. The fix keeps getBinary correct on its own instead of relying on that order.

Rollback safety (please read)

Enabling the flag from this PR on is one-way for content written while it is on: check-in stores the binary only under a revision key that neither an older release nor this release with the flag off reads. Leaving the flag off, the default, is safe to roll back. The doc's "Enabling the flag is not rollback-safe" section says the same. This PR should carry the rollback-unsafe label.

Deliberately not in this PR

Temporary-upload metadata (5), backfill and starter import/export (3), publishing (4), renditions (6). BinaryCleanupJob changes entirely in 5.

Checklist

  • Tests: the 145 unit tests in the doc's run command pass after the review fixes against disposable MinIO and PostgreSQL (1 skip, STS). Integration on this branch before the review fixes, flag off: 248 run, 0 failures, across ContentletAPITest (183), FileMetadataAPITest (48), MetadataDelegateTest and the three new classes; one index-count case (testRemoveContentFromIndexMultilingualContent) failed once and passed on rerun, in two separate flag-off runs. Flag on against MinIO: 17 run, 0 failures, 0 skips. After the review fixes, the integration suites were run on the top of the stack (feat(storage): serve renditions, compiled CSS and templates through S3 asset storage #37776), flag off and on. The three new integration classes are registered in Junit5Suite1; their S3 cases skip in CI, which runs flag-off. After the second fix, flag on against MinIO: BinaryAssetStorageIntegrationTest 10 run, 0 failures, and its new case fails without the fix; BinaryAssetReferenceTest 7 run, 0 failures.
  • Translations: not applicable.
  • Security Implications Contemplated: both cleanup jobs reject direct submissions through the public job endpoint; only the deletion APIs create them. Recovery ZIPs contain content data and live in the configured bucket, so bucket access controls apply to them.

@claude

claude Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @jcastro-dotcms's task in 3m 1s —— View job


Code review

Reviewed against origin/s3-stack/1b-binary-api. All behavior is gated behind FEATURE_FLAG_S3_ASSET_STORAGE (default off); findings below are flag-on only. I rechecked the open dotbot/semgrep findings and looked for new issues.

New Issues

None beyond the already-reported items below.

Existing (still present)

  • 🟡 Medium: dotCMS/src/main/java/com/dotcms/storage/binary/BinaryFieldCleanupProcessor.java:115 and :183 — Postgres-only SQL. Line 183 hardcodes contentlet_as_json = ?::jsonb; line 115 uses parameters || ?::jsonb, parameters->>'afterInode' and returning id. The same PR already branches on DbConnectionFactory.isPostgres() for the identical write in ESContentletAPIImpl.java:6183 (?::jsonb vs ?), so the intent is portability — these two spots are inconsistent with that and would make binaryFieldCleanup jobs fail at runtime (and retry forever) on MySQL/MSSQL/Oracle with the flag on. Either make both portable the same way, or if the S3 lifecycle is deliberately Postgres-only, gate the flag on isPostgres() so it fails fast at enablement rather than mid-cleanup. Fix this →

  • 🟡 Medium: dotCMS/src/main/java/com/dotcms/util/marshal/JacksonMarshalUtilsImpl.java:27 — flag-on registration of ParameterNamesModule on the shared defaultMapper changes constructor selection for every MarshalUtils caller (push-publish bundles, system events, workflow payloads), not just binary field immutables. Also, defaultMapper is Lazy, so the mapper's config is frozen to the flag's value at first use — toggling the flag afterward has no effect, and the behavior differs depending on what triggered the first marshal. Scope the module to a dedicated mapper used only for the binary immutables.

  • 🟡 Medium: dotCMS/src/main/java/com/dotcms/content/elasticsearch/business/ESContentletAPIImpl.java:6181 — with the flag on, every binary-content check-in writes contentlet_as_json unconditionally, even when SAVE_CONTENTLET_AS_JSON is disabled, silently repopulating a column an operator deliberately keeps null. Either honor the property, or fail fast / document that enabling the flag implies JSON persistence.

  • 🟡 Medium: dotCMS/src/main/java/com/dotcms/storage/binary/ContentletBackupStorage.java:69-73 — when contentlet_as_json is blank, ContentletAPI.find(...) can return null; new Contentlet(null) (or the subsequent validateId(content.getIdentifier())) then NPEs inside the deletion transaction, aborting the delete. Narrow (flag on + blank JSON, which the flag normally prevents), but worth a null check with a clear DotDataException. Fix this →

Resolved / not a defect

  • ✅ ESContentletAPIImpl.java:2749 — the "commit listener registered after LocalTransaction already committed" concern does not hold. HibernateUtil.addCommitListener falls back to listener.run() immediately when there is no active transaction (HibernateUtil.java:1060-1062). For a non-transactional caller the flag-on path (LocalTransaction.wrap commits, then not-in-tx → listener runs now) matches the flag-off path (deleteContentlets is not @WrapInTransaction, so the listener also runs immediately). For a transactional caller, LocalTransaction.wrap joins the outer tx and the listener registers to it as before. ContentletDeletedEvent fires in both cases — no new bug.

  • ✅ BinaryFieldCleanupProcessor.java:101,111,117 (semgrep CUSTOM_INJECTION-2) — false positives. All variable values go through addParam(...) placeholders; only fixed fragments are concatenated. No attacker-controlled input reaches SQL syntax.

Notes

  • The rollback-unsafety (H-5/H-1) raised repeatedly by the bot is by design and documented (docs/testing/BINARY_S3_STORAGE.md), and the PR already carries the rollback-unsafe label. Nothing to change in code; just ensure release notes flag that enabling FEATURE_FLAG_S3_ASSET_STORAGE is one-way.

None of the above block the stacked rollout — all are flag-on only (default off), with the Postgres-only SQL being the most actionable since it diverges from the PR's own portable write.
· s3-stack/2-content

@claude

claude Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Pull Request Unsafe to Rollback!!!

  • Category: H-5 — Binary Storage Provider Change
  • Risk Level: 🟠 HIGH
  • Why it's unsafe: This PR extends the opt-in S3 asset lifecycle (FEATURE_FLAG_S3_ASSET_STORAGE, default off) so that, once enabled, new binary check-ins are written under revision-keyed paths (.revisions//...) and referenced only via new storageKey/metadataStorageKey attributes added to the Binary field's contentlet_as_json entry — no legacy flat file is written anymore, and the local copy can be evicted to S3. N-1 (and this same release with the flag turned back off) has no reader for storageKey/metadataStorageKey; it falls back to the legacy flat-file lookup, which finds nothing for content written while the flag was on. The PR's own documentation explicitly calls this out as a one-way step. Additionally, BinaryAssetCleanupProcessor/BinaryFieldCleanupProcessor physically delete local/S3 bytes for deleted content/fields while the flag is on (keeping only a recovery ZIP in a separate S3 deleted-content-backups group) — after a rollback, N-1 cannot reconstruct those bytes on its own; recovery requires manually pulling the ZIP via ContentletBackupStorage.
  • Code that makes it unsafe:
    • docs/testing/BINARY_S3_STORAGE.md — new section "Enabling the flag is not rollback-safe": "Enabling the flag is a one-way step for any content written while it is on... Neither a release without this code nor this release with the flag turned back off reads those revision keys... treat enabling the flag as requiring a forward-only recovery plan."
    • dotCMS/src/main/java/com/dotcms/content/elasticsearch/business/ESContentletAPIImpl.java (handleBinaries, ~line 6990+) — when the flag is enabled, binaries are written via binaryStorageAPI.storeRevision(newInode, velocityVarNm, oldFileName, incomingFile) into the new revision-keyed layout instead of the legacy flat file.
    • dotCMS/src/main/java/com/dotcms/contenttype/model/type/system/AbstractBinaryFieldType.java — adds storageKey()/metadataStorageKey() (@JsonInclude(NON_NULL)) to the Binary field's JSON model, consumed only by flag-aware readers.
    • dotCMS/src/main/java/com/dotcms/storage/binary/BinaryAssetCleanupProcessor.java and BinaryFieldCleanupProcessor.java — delete source binaries/metadata from local disk and S3 once a job completes, leaving only a recovery archive (ContentletBackupStorage) for manual retrieval.
  • Alternative (if possible): This already follows the safest practical pattern for H-5 by shipping the change behind a flag that defaults to off (unchanged behavior for all current deployments). The remaining gap is only for operators who actively enable the flag: document in release notes that flipping FEATURE_FLAG_S3_ASSET_STORAGE on is a point of no return for rollback of any content touched afterward, and/or consider a transitional mode that also writes the legacy flat-file copy while the flag is on, so N-1 can still resolve binaries after a rollback (per the "read from both old and new locations" pattern in the reference doc).

Comment on lines +115 to +117
final var saved = new DotConnect().setSQL("update job set parameters = parameters || ?::jsonb, "
+ "updated_at = current_timestamp where id = ? and queue_name = ? "
+ "and coalesce(parameters->>'afterInode', '') = ? returning id")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Medium severity blocking issue identified in your code:
The method identified is susceptible to injection. The input should be validated and properly
escaped.

Why this might be safe to ignore:

The SQL is assembled only from fixed application-controlled fragments, while all variable values are supplied through parameter binding with placeholders. No attacker-controlled input reaches the SQL syntax, so this is a false positive.

To resolve this comment:

🔧 No guidance has been designated for this issue. Fix according to your organization's approved methods.

💬 Ignore this finding

Reply with Semgrep commands to ignore this finding.

  • /fp <comment> for false positive
  • /ar <comment> for acceptable risk
  • /other <comment> for all other reasons

Alternatively, triage in Semgrep AppSec Platform to ignore the finding created by CUSTOM_INJECTION-2.

If this is a critical or high severity finding, please also link this issue in the #security channel in Slack.

You can view more details about this finding in the Semgrep AppSec Platform.

Comment on lines +110 to +111
final var rows = new DotConnect().setSQL("select inode, identifier, contentlet_as_json "
+ "from contentlet where inode = ? and mod_date <= ? for update")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Medium severity blocking issue identified in your code:
The method identified is susceptible to injection. The input should be validated and properly
escaped.

Why this might be safe to ignore:

The SQL is assembled from fixed query fragments, and the inode and date values are bound through placeholders with addParam rather than interpolated into the statement. No attacker-controlled input reaches SQL syntax in this code.

To resolve this comment:

🔧 No guidance has been designated for this issue. Fix according to your organization's approved methods.

💬 Ignore this finding

Reply with Semgrep commands to ignore this finding.

  • /fp <comment> for false positive
  • /ar <comment> for acceptable risk
  • /other <comment> for all other reasons

Alternatively, triage in Semgrep AppSec Platform to ignore the finding created by CUSTOM_INJECTION-2.

If this is a critical or high severity finding, please also link this issue in the #security channel in Slack.

You can view more details about this finding in the Semgrep AppSec Platform.

Comment on lines +100 to +101
final var candidates = new DotConnect().setSQL("select inode from contentlet where structure_inode = ? "
+ "and inode > ? and mod_date <= ? order by inode limit 100")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Medium severity blocking issue identified in your code:
The method identified is susceptible to injection. The input should be validated and properly
escaped.

Why this might be safe to ignore:

The SQL is built only from fixed query text, while type, cursor, and date values are supplied separately through placeholders and addParam. No attacker-controlled input is interpolated into the SQL syntax.

To resolve this comment:

🔧 No guidance has been designated for this issue. Fix according to your organization's approved methods.

💬 Ignore this finding

Reply with Semgrep commands to ignore this finding.

  • /fp <comment> for false positive
  • /ar <comment> for acceptable risk
  • /other <comment> for all other reasons

Alternatively, triage in Semgrep AppSec Platform to ignore the finding created by CUSTOM_INJECTION-2.

If this is a critical or high severity finding, please also link this issue in the #security channel in Slack.

You can view more details about this finding in the Semgrep AppSec Platform.

@semgrep-dotcms

Copy link
Copy Markdown
Contributor

Semgrep found 1 CUSTOM_INJECTION-2 finding:

  • dotCMS/src/main/java/com/dotcms/storage/binary/BinaryFieldCleanupProcessor.java

🟡 Medium severity issue identified in your code:

The method identified is susceptible to injection. The input should be validated and properly
escaped.

If this is a critical or high severity finding, please also link this issue in the #security channel in Slack.

@claude

claude Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Pull Request Unsafe to Rollback!!!

  • Category: H-5 — Binary Storage Provider Change
  • Risk Level: 🟠 HIGH
  • Why it's unsafe: This PR wires binary Content fields into the pluggable StoragePersistenceAPI/S3 provider when FEATURE_FLAG_S3_ASSET_STORAGE is enabled. With the flag on, check-in stores the active binary only under a .revisions/<uuid>/ S3 key recorded in contentlet_as_json as new storageKey/metadataStorageKey properties — no legacy flat file is written on disk. Neither an N-1 binary (which has no concept of storageKey at all) nor this same release with the flag toggled back off can resolve those revision keys; both fall back to the legacy on-disk field-folder lookup, so any binary written while the flag was on resolves as missing or stale after rollback. The PR's own testing doc states this plainly: "Enabling the flag is not rollback-safe... Keeping a legacy-path copy for rollback is not implemented; treat enabling the flag as requiring a forward-only recovery plan." Blast radius is scoped to deployments that actively opt into the flag (default is disabled, which "changes nothing"), but for any environment that enables it, this is a real H-5 match, not a hypothetical one — it's explicitly what the PR documents.
  • Code that makes it unsafe:
    • docs/testing/BINARY_S3_STORAGE.md (new "### Enabling the flag is not rollback-safe" section) — the PR authors' own rollback-unsafety disclosure.
    • dotCMS/src/main/java/com/dotcms/content/model/type/system/AbstractBinaryFieldType.java — adds storageKey()/metadataStorageKey() to the binary field JSON model (additive, no CURRENT_MODEL_VERSION bump, but still unreadable by any binary lacking the S3 code path).
    • dotCMS/src/main/java/com/dotcms/contenttype/model/field/BinaryField.java — getFieldValue() sets storageKey/metadataStorageKey when AssetStorageFeature.isEnabled().
    • dotCMS/src/main/java/com/dotmarketing/portlets/contentlet/model/Contentlet.java — getBinary() branches entirely on AssetStorageFeature.isEnabled(), resolving via BinaryAssetStorageAPI.getRevisionFile()/getBinaryFile() instead of the legacy field folder when enabled.
    • dotCMS/src/main/java/com/dotcms/business/json/ContentletJsonAPIImpl.java getBinary() — same enabled/disabled fork; the disabled/legacy path cannot see revision-keyed binaries.
    • dotCMS/src/main/java/com/dotcms/content/elasticsearch/business/ESContentletAPIImpl.java (checkin, handleBinaries) — persists contentlet_as_json with the new storage keys and reroutes binary handling through BinaryAssetStorageAPI only when the flag is on.
  • Alternative (if possible): The safer H-5 pattern is a fallback chain — keep writing (or reading through) the legacy flat file/local path in addition to the S3-backed revision key while the flag is enabled, so a rollback or flag disablement still finds assets at the classic path, until the old path is retired in a later release. The PR's own doc confirms this dual-path fallback is not implemented today.

@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Pull Request Unsafe to Rollback!!!

  • Category: H-5 — Binary Storage Provider Change
  • Risk Level: 🟠 HIGH
  • Why it's unsafe: The PR adds an opt-in S3 binary-asset storage lifecycle behind AssetStorageFeature.FLAG = "FEATURE_FLAG_S3_ASSET_STORAGE" (default off, so a default deployment is unaffected). But once an operator turns the flag on, new/edited binary field checkins are written under an immutable revision key (binary-assets/{a}/{b}/{inode}/{field}/.revisions/{revision-id}/{filename}) recorded only in contentlet_as_json (storageKey/metadataStorageKey) — no legacy flat file is written, and the local copy can later be evicted to S3. N-1 (and this same release with the flag turned back off) never reads those revision keys; its binary resolution falls back to the legacy field folder, so affected binaries resolve as stale or missing after rollback. This is exactly the scenario category H-5 describes ("N-1 after rollback reverts to the File System provider and looks for files at the old filesystem path — those files do not exist there").
  • Code that makes it unsafe:
    • dotCMS/src/main/java/com/dotcms/storage/AssetStorageFeature.java:18 — the flag itself, default-off gate (FLAG, isEnabled()).
    • dotCMS/src/main/java/com/dotcms/content/business/json/ContentletJsonAPIImpl.java (getBinary, hydrate... methods around lines 401-437, 494-503) — reads/writes storageKey/metadataStorageKey in contentlet_as_json and resolves binaries through BinaryAssetReference/BinaryAssetStorageAPI only when the flag is enabled; a flag-off reader has no path back to those keys.
    • dotCMS/src/main/java/com/dotcms/content/elasticsearch/business/ESContentletAPIImpl.java (handleBinaries, ~lines 6835-7000) — under the flag, checkin calls binaryStorageAPI.storeRevision(...) instead of writing the legacy filesystem path, explicitly commented "Never overwrite bytes referenced by committed content. The JSON reference switches in the database transaction; rollback leaves the old object intact" (i.e., the old revision survives, but a rolled-back N-1 binary can no longer find the new revision it never wrote to the legacy path).
    • docs/testing/BINARY_S3_STORAGE.md — new section "### Enabling the flag is not rollback-safe" added by this PR, which states verbatim: "Enabling the flag is a one-way step for any content written while it is on... Neither a release without this code nor this release with the flag turned back off reads those revision keys..., so affected binaries resolve as stale or missing... treat enabling the flag as requiring a forward-only recovery plan." The PR author has already self-identified this exact risk.
  • Alternative (if possible): This already follows the safer alternative reasonably well — the change is fully opt-in and off by default, and content written before the flag was ever enabled is untouched by rollback. The residual risk is operational, not code-level: document (in release notes / ops runbook) that toggling FEATURE_FLAG_S3_ASSET_STORAGE on in production is a one-way gate per the two-phase-migration spirit of C-1/H-5 — i.e., do not enable it until the previous release is fully retired from the rollback window, and never toggle it back off as a "rollback" step once binaries have been written under it.

Pull Request Unsafe to Rollback!!!

  • Category: H-1 — One-Way Data Migration or Destructive Backfill
  • Risk Level: 🟠 HIGH
  • Why it's unsafe: When the same feature flag is enabled, field/content deletion permanently destroys binary data via background jobs rather than the old synchronous filesystem delete. BinaryFieldCleanupProcessor.cleanInode (new file) deletes S3 metadata and source objects and rewrites contentlet_as_json to drop the field, and BinaryAssetCleanupProcessor/DropOldContentletRunner queue full-inode binary deletion. These are one-way destructive operations (consistent with H-1's "DELETE of business records" / "TRUNCATE + rebuild" signal) that, combined with H-5's storage-path switch, leave no rollback-readable copy once cleanup runs.
  • Code that makes it unsafe:
    • dotCMS/src/main/java/com/dotcms/storage/binary/BinaryFieldCleanupProcessor.java:230-260 (cleanInode) — deletes metadata objects, then source binaries (binariesAPI.deleteBinaryPaths, binariesAPI.deleteGeneratedFiles) and the legacy cache directory, after rewriting contentlet_as_json to remove the field in archiveRow (same file, ~lines 165-190).
    • dotCMS/src/main/java/com/dotcms/content/elasticsearch/business/DropOldContentletRunner.java:176-181, 232-252 — under the flag, enqueues BinaryAssetCleanupProcessor.enqueue(inode) and calls APILocator.getBinaryAssetStorageAPI().deleteAllBinaries(inode) instead of the reversible legacy deleteFromAssetsDir.
  • Alternative (if possible): The PR already mitigates this with a recovery path — ContentletBackupStorage writes a verified S3 recovery ZIP before destructive deletion when BACKUP_DELETED_CONTENTLETS_TO_DISK is also enabled, and BinaryFieldCleanupProcessor.archiveRow archives the exact inventory before cleanInode deletes it. Recommend making the backup-before-destroy path mandatory (not conditional on a second flag) whenever FEATURE_FLAG_S3_ASSET_STORAGE is on, and calling out in release notes that recovery after cleanup is manual (S3 ZIP download), not an automatic rollback.

Both findings are conditional on FEATURE_FLAG_S3_ASSET_STORAGE being turned on (default: off) — with the flag off, this PR changes nothing and rollback is unaffected, per the PR's own docs/testing/BINARY_S3_STORAGE.md.

@claude

claude Bot commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

test connectivity check - ignore

@claude

claude Bot commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Pull Request Unsafe to Rollback!!!

  • Category: H-5 — Storage Provider Configuration Change

  • Risk Level: 🟠 HIGH

  • Why it's unsafe: This PR builds out an opt-in S3 binary-asset lifecycle (AssetStorageFeature.FLAG = "FEATURE_FLAG_S3_ASSET_STORAGE", default false). When a site turns the flag on, ESContentletAPIImpl.handleBinaries() starts writing new binaries through BinaryAssetStorageAPI.storeRevision(...) (S3) instead of the legacy filesystem path, and Contentlet.getBinary() / ContentletJsonAPIImpl.getBinary() resolve files via the new storageKey/metadataStorageKey references in contentlet_as_json rather than the on-disk {inode}/{field}/ folder. N-1 has no code for BinaryAssetStorageAPI, BinaryAssetReference, or the storageKey/metadataStorageKey JSON properties at all — it will silently fall back to the legacy filesystem lookup and find nothing for any content checked in (or whose binary field was edited) while the flag was on, exactly the H-5 failure mode ("N-1 reverts to File System provider and looks for files at the old path — those files do not exist there").

  • Code that makes it unsafe:

    • dotCMS/src/main/java/com/dotcms/storage/AssetStorageFeature.java (flag + default false)
    • dotCMS/src/main/java/com/dotmarketing/portlets/contentlet/model/Contentlet.java — getBinary(String) branches on AssetStorageFeature.isEnabled() and resolves through BinaryAssetStorageAPI.getBinaryFile/getRevisionFile instead of the filesystem folder
    • dotCMS/src/main/java/com/dotcms/content/business/json/ContentletJsonAPIImpl.java — getBinary(Field, String, FieldValue) and setValue(...) read/write the new storageKey/metadataStorageKey fields
    • dotCMS/src/main/java/com/dotcms/content/elasticsearch/business/ESContentletAPIImpl.java — handleBinaries(...) writes via binaryStorageAPI.storeRevision(newInode, velocityVarNm, oldFileName, incomingFile) when the flag is enabled
  • Alternative (if possible): This is already built as an opt-in, additive flag with the JSON fields marked @JsonInclude(NON_NULL) so disabled installs are unaffected — that part follows the two-phase pattern correctly. The remaining gap is operational: document in the release notes that flipping FEATURE_FLAG_S3_ASSET_STORAGE on is a one-way move per the safer-alternative note in H-5 ("treat it as infrastructure configuration"), and that a rollback to N-1 after enabling it requires restoring binaries from the S3 group back to the filesystem path before N-1 can serve them again.

  • Category: H-1 — One-Way Data Migration or Destructive Backfill

  • Risk Level: 🟠 HIGH

  • Why it's unsafe: When the same flag is enabled, BinaryFieldCleanupProcessor.archiveRow(...) removes a binary field entirely from contentlet_as_json (((ObjectNode) document.path("fields")).remove(field)) and BinaryAssetCleanupProcessor.process(...) / BinaryFieldCleanupProcessor.cleanInode(...) permanently delete the field's local files, generated files, legacy cache and stored metadata. The only recovery path is a proprietary ZIP archive written by the new ContentletBackupStorage class (contentlet.json + contentlet.xml + assets/... entries) — a format N-1 has no reader for. After rollback, N-1 cannot reconstruct the removed field data even though a "backup" technically exists, because N-1 lacks ContentletBackupStorage.open(...) entirely.

  • Code that makes it unsafe:

    • dotCMS/src/main/java/com/dotcms/storage/binary/BinaryFieldCleanupProcessor.java — archiveRow(...) strips the field from the JSON and calls ContentletBackupStorage.getInstance().storeField(...); cleanInode(...) then deletes the archived binaries/metadata for good
    • dotCMS/src/main/java/com/dotcms/storage/binary/BinaryAssetCleanupProcessor.java — process(Job) calls fileMetadataAPI.removeMetadataForInode(...) and binaries.deleteAllBinaries(inode) once the owning contentlet row is gone
    • dotCMS/src/main/java/com/dotcms/storage/binary/ContentletBackupStorage.java — new ZIP-based recovery format (store, storeField, open) with no equivalent on N-1
  • Alternative (if possible): Per H-1's safer alternative, prefer additive-only removal: keep the field/binary data in place (or in a format N-1 already understands, e.g. the existing filesystem cache) for at least one release cycle after the flag is enabled, only physically deleting once N-1 is fully outside the supported rollback window.

Note: both risks are conditional on FEATURE_FLAG_S3_ASSET_STORAGE actually being turned on — with the flag at its default (false), all legacy filesystem code paths are preserved unchanged and this PR is rollback-safe on its own. Flag the risk in release notes for any environment that enables S3 asset storage.

@swicken
swicken force-pushed the s3-stack/2-content branch from ea2656b to f874dda Compare October 2, 2026 16:31
@swicken
swicken force-pushed the s3-stack/1b-binary-api branch from 666a133 to 7e8397b Compare October 2, 2026 16:31
@claude

claude Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Pull Request Unsafe to Rollback!!!

  • Category: H-5 — Binary Storage Provider Change
  • Risk Level: 🟠 HIGH
  • Why it's unsafe: This PR adds a new, opt-in S3-backed binary asset lifecycle gated by AssetStorageFeature.isEnabled() (default false, read from FEATURE_FLAG_S3_ASSET_STORAGE). When the flag is turned on in an environment, binaries are no longer written to the legacy filesystem path during check-in — they are uploaded as immutable revisions via BinaryAssetStorageAPI.storeRevision(...) and referenced by a new storageKey/metadataStorageKey pair stored on the binary field. Legacy on-disk copies and the legacy image-rendition cache are then actively deleted by two new background job processors (BinaryAssetCleanupProcessor, BinaryFieldCleanupProcessor) once content is deleted or a field is removed. N-1 has no knowledge of AssetStorageFeature, storageKey, or the S3-backed BinaryAssetStorageAPI — it always falls back to the legacy filesystem path (ESContentletAPIImpl.getBinaryFile, Contentlet.getBinary, FileAsset.getFileAsset, BinaryExporterServlet.doGet). After a rollback to N-1, any binary that was created, replaced, or whose legacy cache was reclaimed while the flag was on will not be found on the local filesystem — file downloads, image rendering, and metadata reads for that content return 404s or throw, exactly as described in the reference doc's H-5 example ("File System → S3" switch).
  • Code that makes it unsafe:
    • dotCMS/src/main/java/com/dotcms/storage/AssetStorageFeature.java — new opt-in flag FEATURE_FLAG_S3_ASSET_STORAGE controlling the storage provider switch.
    • dotCMS/src/main/java/com/dotcms/content/elasticsearch/business/ESContentletAPIImpl.java — handleBinaries(...) now calls binaryStorageAPI.storeRevision(newInode, velocityVarNm, oldFileName, incomingFile) instead of writing to the local filesystem when the flag is enabled; getBinaryFile(...) reads exclusively from APILocator.getBinaryAssetStorageAPI() instead of the legacy path; deleteBinaryFiles(...) enqueues BinaryAssetCleanupProcessor.enqueue(...) which deletes the legacy local cache folder (FileUtil.deltree(legacyCache)).
    • dotCMS/src/main/java/com/dotcms/storage/binary/BinaryAssetCleanupProcessor.java (new file) — process(Job) deletes recorded binary paths and the legacy filesystem image-rendition cache (new File(..., "cache/" + inode.charAt(0) + "/" + inode.charAt(1) + "/" + inode)), making the legacy path unrecoverable by N-1 after rollback.
    • dotCMS/src/main/java/com/dotcms/storage/binary/BinaryFieldCleanupProcessor.java (new file) — archives/removes binary field data per content type via the same S3-lifecycle queue.
    • dotCMS/src/main/java/com/dotmarketing/portlets/contentlet/model/Contentlet.java — getBinary(...) resolves through APILocator.getBinaryAssetStorageAPI().getBinaryFile(...) / .getRevisionFile(...) instead of the legacy inode-keyed filesystem folder when the flag is on.
    • dotCMS/src/main/java/com/dotcms/content/model/type/system/AbstractBinaryFieldType.java and dotCMS/src/main/java/com/dotcms/contenttype/model/field/BinaryField.java — new storageKey/metadataStorageKey properties that encode the S3 location; unrecognized/unusable by N-1's filesystem-only binary resolution.
  • Alternative (if possible): Per the reference doc's safer alternative for H-5 — "support reading from both old and new locations (fallback chain) before removing the old path." This PR's cleanup processors should not delete the legacy filesystem copy (or legacy cache) at all for at least one release cycle, so N-1 can still resolve binaries from disk after a rollback while N reads from S3 first. Alternatively, since the flag already defaults to disabled, explicitly document in release notes that enabling FEATURE_FLAG_S3_ASSET_STORAGE in production is a one-way, rollback-unsafe operational change, and treat flipping it on (not just the code shipping) as the unsafe event.

@swicken
swicken marked this pull request as ready for review October 2, 2026 18:24
@nollymar nollymar added the PR : dotbot review Trigger dotbot AI code review and the post-merge QA test plan label Oct 6, 2026
final String backup = ContentletBackupStorage.getInstance().storeField(
row.get("identifier").toString(), inode, json, binaries, List.copyOf(metadata));
((ObjectNode) document.path("fields")).remove(field);
new DotConnect().setSQL("update contentlet set contentlet_as_json = ?::jsonb where inode = ?")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 [P1] BinaryFieldCleanupProcessor.java:183 make JSON update portable

Current code:

new DotConnect().setSQL("update contentlet set contentlet_as_json = ?::jsonb where inode = ?")
        .addParam(JSON.writeValueAsString(document)).addParam(inode).loadResult();

Problem: Hardcoded ::jsonb cast fails on MySQL, MSSQL and Oracle.

Fix:

new DotConnect().setSQL("update contentlet set contentlet_as_json = "
        + (DbConnectionFactory.isPostgres() ? "?::jsonb" : "?") + " where inode = ?")
        .addParam(JSON.writeValueAsString(document)).addParam(inode).loadResult();


private final Lazy<ObjectMapper> defaultMapper = Lazy.of(() -> {
final ObjectMapper objectMapper = new ObjectMapper();
if (com.dotcms.storage.AssetStorageFeature.isEnabled()) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 [P2] JacksonMarshalUtilsImpl.java:27 flag-on ParameterNamesModule alters the shared global mapper

Current code:

if (com.dotcms.storage.AssetStorageFeature.isEnabled()) {
    objectMapper.registerModule(new com.fasterxml.jackson.module.paramnames.ParameterNamesModule(
            com.fasterxml.jackson.annotation.JsonCreator.Mode.PROPERTIES));
}

Problem: Flag-on registration changes constructor selection for every MarshalUtils caller (push-publish bundles, system events, workflow payloads), not just binary types. Mapper behavior also depends on the flag's value at first lazy use.

Fix:

// Scope ParameterNamesModule to a dedicated mapper used only for the binary field
// immutables, leaving the shared defaultMapper configuration unchanged.

…3 asset storage

Third slice of the S3 asset storage work. With FEATURE_FLAG_S3_ASSET_STORAGE on,
content check-in stores immutable binary revisions and records their storage keys
in the content JSON, metadata generation restores evicted originals and shares
byte-derived extraction, and deletion uses durable binaryAssetCleanup and
binaryFieldCleanup jobs plus verified S3 recovery archives. The S3 cleanup
processors do not register while the flag is off, and flag-off behavior matches main.
Field-removal archiving used to run as one transaction over every version of a
content type, holding FOR UPDATE locks on each batch while it uploaded recovery
ZIPs to S3, so large types blocked edits and risked timeouts, and one failure
rolled back and repeated all the work. Each row is now archived in its own short
transaction that locks only that row, rechecks the deletion timestamp, and
commits a compare-and-set afterInode cursor with it, so a retry resumes after the
last archived row. ContentletAPI.cleanField queues the same job instead of
uploading inline. Flag-off behavior is unchanged.
…and per-language binaries

Whole-inode cleanup now records the inode's exact binary and metadata paths when the deletion
enqueues the job, and the worker deletes only those paths. An inode re-created after deletion
(push publishing keeps the sender's inodes) keeps the revisions it uploads. The "content version
still exists" refusal is unchanged, and a job without a recorded inventory deletes nothing.
FileMetadataAPI.removeMetadataForInode is split into listMetadataForInode and removeMetadataPaths
for this.

BinaryAssetCleanupProcessor now implements Validator and rejects submissions through the public
job endpoint, as BinaryFieldCleanupProcessor already does.

A check-in that uploads a new revision registers a rollback listener that deletes that revision
and its revision metadata. The key carries a fresh UUID, so no other version can reference it. A
failed delete is logged and never thrown.

With the flag on, deleting one language of multilingual content now records recovery archives
and cleanup jobs like the other deletion paths, and ContentletAPI.delete(Contentlet) runs in a
transaction so those jobs commit with the deleted rows (Site Browser and WebDAV call it without
one). The flag-off paths are unchanged.

deleteBinaryFiles enqueues one cleanup job per distinct inode, since callers pass duplicated
version lists.

BinaryExporterServlet releases the cache lease once the file to stream is open, so a slow client
no longer defers eviction.

DropOldContentletRunner.deleteFromAssetsDir is restored to main; its flag-on branch was
unreachable.

The S3 storage doc is updated to match, and the unit and flag-on integration tests cover the
exact inventory, the endpoint rejection, the rollback deletion, the per-language path and the
single job per inode.

The S3 metadata outage test now expects an unreadable local copy to be replaced from S3 when S3
holds a readable one, and to fail without deleting it when it does not, matching the storage
layer.

The S3 servlet response test now expects eviction to proceed once the full response has opened
its file, and to be refused in the range branch until that branch opens it.
…r its key names

Contentlet.getBinary resolved a revision key against the contentlet's current
inode. A contentlet that carries the previous version's revision under a new
inode (as check-in does after it assigns the new inode) failed with
"File System error." once that revision was evicted from the local cache,
because the key belongs to the old inode. getBinary now restores the revision
under the owner its key names; BinaryAssetReference.localFile still validates
the key.

Check-in does not hit this today only because saving the content JSON restores
the file while hydrating metadata, before check-in reads the binary.

Refs #37868
@swicken
swicken force-pushed the s3-stack/2-content branch from 3523198 to 236a0d2 Compare October 7, 2026 20:47
@swicken
swicken force-pushed the s3-stack/1b-binary-api branch from 7e8397b to 14d49ce Compare October 7, 2026 20:47
final String backup = ContentletBackupStorage.getInstance().storeField(
row.get("identifier").toString(), inode, json, binaries, List.copyOf(metadata));
((ObjectNode) document.path("fields")).remove(field);
new DotConnect().setSQL("update contentlet set contentlet_as_json = ?::jsonb where inode = ?")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 [P1] BinaryFieldCleanupProcessor.java:183 make JSON update portable

Current code:

new DotConnect().setSQL("update contentlet set contentlet_as_json = ?::jsonb where inode = ?")

Problem: Hardcoded ::jsonb cast fails on MySQL, MSSQL and Oracle.

Fix:

new DotConnect().setSQL("update contentlet set contentlet_as_json = "
        + (DbConnectionFactory.isPostgres() ? "?::jsonb" : "?") + " where inode = ?")

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

dotbot code review:

  • Reviewer: meta/muse-spark-1.3 (medium)
  • Overall: patch is incorrect
  • New findings this run: 1
  • Prior unresolved dotbot findings still relevant: 0
  • Active findings total: 1

Flag-on field cleanup hardcodes Postgres-only ::jsonb for contentlet_as_json while the same patch branches on isPostgres() elsewhere, breaking cleanup on other supported databases.

Tip: comment with "/dotbot address comments" to attempt automated fixes for unresolved review threads.

reviewed by dotbot · meta/muse-spark-1.3 · medium

for (String path : storage.listObjectPaths(metadataGroup(), parent)) {
final String absolute = path.startsWith("/") ? path : "/" + path;
if (parent.equals(legacyParent) && absolute.substring(parent.length()).contains("/")) continue;
if (ownsMetadata(owner, absolute)) metadata.add(absolute);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 [P2] BinaryFieldCleanupProcessor.java:183 cursor update uses Postgres-only JSONB concatenation

Current code:

final var saved = new DotConnect().setSQL("update job set parameters = parameters || ?::jsonb, "
                + "updated_at = current_timestamp where id = ? and queue_name = ? "
                + "and coalesce(parameters->>'afterInode', '') = ? returning id")

Problem: || ?::jsonb concatenation, parameters->>'...' extraction and returning id are Postgres-only; this job fails at runtime on MySQL/MSSQL/Oracle even before the archived row update.

Fix:

final var saved = new DotConnect().setSQL("update job set parameters = "
        + (DbConnectionFactory.isPostgres() ? "parameters || ?::jsonb" : "?")
        + ", updated_at = current_timestamp where id = ? and queue_name = ? "
        + (DbConnectionFactory.isPostgres() ? "and coalesce(parameters->>'afterInode', '') = ? returning id" : "...")

The cursor compare/merge needs per-DB SQL (e.g., MySQL JSON_MERGE_PATCH(parameters, ?) and JSON_UNQUOTE(JSON_EXTRACT(parameters, '$.afterInode')), plus a follow-up select for rowcount). If only Postgres is supported for the S3 lifecycle, gate the flag on isPostgres().

contentletRaw);

if (com.dotcms.storage.AssetStorageFeature.isEnabled() && !contentType.fields(BinaryField.class).isEmpty()) {
// Persist immutable binary references alongside the filename in this transaction.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 [P2] ESContentletAPIImpl.java:6181 unconditionally writes contentlet_as_json, overriding SAVE_CONTENTLET_AS_JSON=false

Current code:

if (com.dotcms.storage.AssetStorageFeature.isEnabled() && !contentType.fields(BinaryField.class).isEmpty()) {
    final String json = APILocator.getContentletJsonAPI().toJson(contentlet);
    final String jsonValue = DbConnectionFactory.isPostgres() ? "?::jsonb" : "?";
    new DotConnect().setSQL("update contentlet set contentlet_as_json = " + jsonValue + " where inode = ?")

Problem: With the flag on, every binary-content checkin writes contentlet_as_json even when the system property SAVE_CONTENTLET_AS_JSON is disabled, silently repopulating a column operators deliberately keep null.

Fix:

if (com.dotcms.storage.AssetStorageFeature.isEnabled()
        && Config.getBooleanProperty("SAVE_CONTENTLET_AS_JSON", true)
        && !contentType.fields(BinaryField.class).isEmpty()) {

If the S3 lifecycle genuinely requires the JSON, document that enabling FEATURE_FLAG_S3_ASSET_STORAGE implies JSON persistence, or fail fast when SAVE_CONTENTLET_AS_JSON=false instead of silently overriding it.

final var rows = new DotConnect().setSQL("select contentlet_as_json from contentlet where inode = ? for update")
.addParam(requested.getInode()).loadObjectResults();
if (rows.isEmpty()) throw new DotDataException("Content version disappeared before backup");
final Object persisted = rows.getFirst().get("contentlet_as_json");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 [P2] ContentletBackupStorage.java:73 identifier of restored contentlet used before existence check

Current code:

final Contentlet content = persisted == null || persisted.toString().isBlank()
        ? new Contentlet(APILocator.getContentletAPI().find(requested.getInode(), APILocator.systemUser(), false))
        : APILocator.getContentletJsonAPI().toMutableContentlet(
                ContentletJsonHelper.INSTANCE.get().immutableFromJson(persisted.toString()));
validateId(content.getIdentifier());

Problem: When contentlet_as_json is blank, ContentletAPI.find returns null or throws after the row was already locked for update; validateId(content.getIdentifier()) on null content NPEs inside a transaction, aborting the deletion.

Fix:

final Contentlet content = persisted == null || persisted.toString().isBlank()
        ? APILocator.getContentletAPI().find(requested.getInode(), APILocator.systemUser(), false)
        : APILocator.getContentletJsonAPI().toMutableContentlet(
                ContentletJsonHelper.INSTANCE.get().immutableFromJson(persisted.toString()));
if (content == null) throw new DotDataException("Content version disappeared before backup");
validateId(content.getIdentifier());

// S3 cleanup jobs must commit with the deleted rows, and callers such as the Site
// Browser and WebDAV reach this method without a transaction.
final boolean[] result = new boolean[1];
LocalTransaction.wrap(() -> result[0] = this.deleteContentlets(contentlets, user,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 [P2] ESContentletAPIImpl.java:2757 commit listener registered after LocalTransaction already committed

Current code:

final boolean[] result = new boolean[1];
LocalTransaction.wrap(() -> result[0] = this.deleteContentlets(contentlets, user,
        respectFrontendRoles, isSite));
deleted = result[0];

Problem: For non-transactional callers (Site Browser/WebDAV), LocalTransaction.wrap commits inside; the ContentletDeletedEvent commit listener is added after commit with no active transaction and never fires.

Fix:

LocalTransaction.wrap(() -> {
    result[0] = this.deleteContentlets(contentlets, user, respectFrontendRoles, isSite);
    HibernateUtil.addCommitListener(() -> this.localSystemEventsAPI.notify(
            new ContentletDeletedEvent<>(contentlet, user)));
});
deleted = result[0];

Assumption: HibernateUtil.addCommitListener no-ops without a current transaction. What to verify: subscribers of ContentletDeletedEvent still run for non-transactional deletes with the flag on.

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

dotbot code review:

  • Reviewer: ~z-ai/glm-latest (medium)
  • Overall: patch is incorrect
  • New findings this run: 4
  • Prior unresolved dotbot findings still relevant: 1
  • Active findings total: 5

4 new actionable findings were identified in the current changes, and 1 prior unresolved dotbot finding still applies, so the patch remains incorrect.

All behavioral changes are gated behind FEATURE_FLAG_S3_ASSET_STORAGE (default off, verified across call sites) and covered by extensive new tests. The found issues (Postgres-only SQL fragments, JSON-column override, backup NPE edge, post-commit listener) are all flag-on and addressable without blocking the stacked rollout, and the flag-on rollback limitation is explicitly documented as intentional.

Tip: comment with "/dotbot address comments" to attempt automated fixes for unresolved review threads.

reviewed by dotbot · ~z-ai/glm-latest · medium

@claude

claude Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Pull Request Unsafe to Rollback!!!

  • Category: H-5 — Binary Storage Provider Change

  • Risk Level: 🟠 HIGH

  • Why it's unsafe: With FEATURE_FLAG_S3_ASSET_STORAGE enabled, ESContentletAPIImpl.handleBinaries() (dotCMS/src/main/java/com/dotcms/content/elasticsearch/business/ESContentletAPIImpl.java, the new com.dotcms.storage.AssetStorageFeature.isEnabled() branch around line 6860) stores every new/replaced binary exclusively through BinaryAssetStorageAPI.storeRevision(...) (S3), instead of copying it into the inode's filesystem folder. N-1's Contentlet.getBinary() (dotCMS/src/main/java/com/dotmarketing/portlets/contentlet/model/Contentlet.java, the pre-existing filesystem branch using binaryFileFolder.listFiles(new BinaryFileFilter())) only knows how to read the old filesystem path. If this release is deployed with the flag on and content is checked in, then rolled back to N-1, N-1 looks for binaries at a filesystem path that was never written — every affected binary field 404s. This is the exact scenario the reference doc's H-5 example describes.

  • Code that makes it unsafe: dotCMS/src/main/java/com/dotcms/content/elasticsearch/business/ESContentletAPIImpl.java (handleBinaries(), new S3-only upload branch using binaryStorageAPI.storeRevision(...)); dotCMS/src/main/java/com/dotcms/storage/AssetStorageFeature.java (the flag gate itself)

  • Alternative (if possible): Keep the flag's default off (already true here) and call this out explicitly in release notes as rollback-unsafe for any environment that has turned it on; longer term, support a fallback read chain (S3 first, then filesystem) so N-1 can still resolve binaries written by N, per the H-5 safer-alternative guidance.

  • Category: H-1 — One-Way Data Migration or Destructive Backfill

  • Risk Level: 🟠 HIGH

  • Why it's unsafe: BinaryFieldCleanupProcessor.archiveRow() (dotCMS/src/main/java/com/dotcms/storage/binary/BinaryFieldCleanupProcessor.java, lines ~376-424) runs when a Binary field is deleted from a content type while the S3 flag is on. It copies the field's bytes/metadata into a private backup ZIP in the deleted-content-backups S3 group, then permanently removes the field from contentlet_as_json for every affected contentlet row: ((ObjectNode) document.path("fields")).remove(field); followed by update contentlet set contentlet_as_json = ... where inode = ?. N-1 has no code path that reads ContentletBackupStorage's backup archives — it only ever reads contentlet_as_json directly. If this release is rolled back after a field deletion has run under the flag, N-1 permanently sees the field's data as gone; recovery requires a human to manually pull and replay the ZIP archive, which is exactly the "no automatic undo" failure mode this category describes.

  • Code that makes it unsafe: dotCMS/src/main/java/com/dotcms/storage/binary/BinaryFieldCleanupProcessor.java (archiveRow(), the document.path("fields") removal + contentlet_as_json update)

  • Alternative (if possible): Follow the additive-migration pattern — leave the field's JSON entry in place (e.g. mark it archived with a pointer to the backup key) rather than removing it outright, so N-1 can still see something for the field; only strip it in a later release once N-1 is fully retired.

Both findings are conditional on operators setting FEATURE_FLAG_S3_ASSET_STORAGE=true (default false, unchanged by this PR) — with the flag off, this PR is behavior-identical to main.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

PR : dotbot review Trigger dotbot AI code review and the post-merge QA test plan

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

3 participants